AIToday
Large Language ModelsOpen-Source AIQiita 機械学習Published: Oct 1, 2026, 13:00 JST

sokudan: Japanese sokudan hits 0.880, beats LayA

sokudan: Japanese sokudan hits 0.880, beats LayA

3 Key Points

  1. What happened

    Developer GeneLab released sokudan, a Japanese-only decision model trained on 314.6M parameters, scoring 0.880 on department routing and 0.844 AUROC on churn in bench_ja, versus laya-multilingual's 0.747 and 0.523.

  2. Why it matters

    This suggests teams handling Japanese business text can run classification locally, keeping data on-premises, and still beat a multilingual baseline, according to the benchmark.

  3. What to watch

    The 0.880 headline figure depends on one-line descriptions added to each choice label; stripping them drops accuracy to 0.537. Watch whether the GitHub issues show the same gap on real Japanese traffic.

WHO IT HITSJapanese-speaking Python developers already calling an LLM API for classification and labeling work are the main audience — they can now swap in a locally run model instead. Teams with data that cannot leave the company are the second group, because the weights are public.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

sokudan did not appear in a vacuum. Its author first measured laya-multilingual, an open multilingual decision model, on 300 Japanese business sentences and found that while department routing worked, the yes/no question about churn hints scored AUROC 0.523 — barely above chance. The first version of sokudan was built over a two-day sprint to fill that gap. It is not the only option: TypeSafe AI released Jev on September 15, 2026, and Liquid AI released d1 on September 29, 2026, but both ship as APIs with closed weights and no published Japanese accuracy. The table in the article lists Jev, d1, Laya, AnyJev from Nokia Applied Research, and sokudan, with sokudan the only one marked both open-weight and Japanese-specific.

The benchmark itself is a 300-item set of Japanese business texts, each labeled with a four-way department choice, a three-level urgency score, and a yes/no churn flag. sokudan scored 0.880 on the choice task, 0.075 RPS on the score task, 0.780 accuracy and 0.844 AUROC on the yes/no task. The author openly notes Jev and d1 are not in the table because TypeSafe AI's terms of service forbid using its service or output to develop similar products, so Jev was never called during development. The article also flags that sokudan's training data is synthetic — 4,833 documents across 21 fields generated by a single LLM — and that bench_ja is synthetic too.

The practical stakes hinge on how much the choice labels matter in the real deployment. The single biggest swing in accuracy came not from model structure or training data but from one line of description per choice: with short labels only, accuracy fell from 0.880 to 0.537. That suggests teams adopting sokudan will need to invest in writing clear option descriptions rather than treating the model as plug-and-play. The author also warns against using it for hiring, credit, disciplinary, medical, or legal decisions, and notes that English inputs perform at majority-vote level.

FAQ
What is a decision model, and how is it different from an LLM?
A decision model reads a text and a typed question and returns one probability per choice in a single pass. Unlike an LLM, it cannot produce an answer outside the given choices, and its probabilities come from the output layer rather than a self-report.
Is sokudan free to use?
Yes. The published weights v0.2 are under Apache-2.0, and the base model is sbintuitions/modernbert-ja-310m under MIT. It installs with pip and supports Python 3.11 to 3.13.
How fast does sokudan run?
On an RTX 5090 it processed bench_ja's 300 items with 3 questions at a median of 22.8 ms per item, using about 1.3 GiB of GPU memory. On a Core Ultra 9 285K CPU the median was 371 ms.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleLASTT sues OpenAI over Hugging Face hack, seeks injunction