
What happened
Developer GeneLab released sokudan, a Japanese-only decision model trained on 314.6M parameters, scoring 0.880 on department routing and 0.844 AUROC on churn in bench_ja, versus laya-multilingual's 0.747 and 0.523.
Why it matters
This suggests teams handling Japanese business text can run classification locally, keeping data on-premises, and still beat a multilingual baseline, according to the benchmark.
What to watch
The 0.880 headline figure depends on one-line descriptions added to each choice label; stripping them drops accuracy to 0.537. Watch whether the GitHub issues show the same gap on real Japanese traffic.
WHO IT HITSJapanese-speaking Python developers already calling an LLM API for classification and labeling work are the main audience — they can now swap in a locally run model instead. Teams with data that cannot leave the company are the second group, because the weights are public.
Summaries like this, in your inbox every morning.
sokudan did not appear in a vacuum. Its author first measured laya-multilingual, an open multilingual decision model, on 300 Japanese business sentences and found that while department routing worked, the yes/no question about churn hints scored AUROC 0.523 — barely above chance. The first version of sokudan was built over a two-day sprint to fill that gap. It is not the only option: TypeSafe AI released Jev on September 15, 2026, and Liquid AI released d1 on September 29, 2026, but both ship as APIs with closed weights and no published Japanese accuracy. The table in the article lists Jev, d1, Laya, AnyJev from Nokia Applied Research, and sokudan, with sokudan the only one marked both open-weight and Japanese-specific.
The benchmark itself is a 300-item set of Japanese business texts, each labeled with a four-way department choice, a three-level urgency score, and a yes/no churn flag. sokudan scored 0.880 on the choice task, 0.075 RPS on the score task, 0.780 accuracy and 0.844 AUROC on the yes/no task. The author openly notes Jev and d1 are not in the table because TypeSafe AI's terms of service forbid using its service or output to develop similar products, so Jev was never called during development. The article also flags that sokudan's training data is synthetic — 4,833 documents across 21 fields generated by a single LLM — and that bench_ja is synthetic too.
The practical stakes hinge on how much the choice labels matter in the real deployment. The single biggest swing in accuracy came not from model structure or training data but from one line of description per choice: with short labels only, accuracy fell from 0.880 to 0.537. That suggests teams adopting sokudan will need to invest in writing clear option descriptions rather than treating the model as plug-and-play. The author also warns against using it for hiring, credit, disciplinary, medical, or legal decisions, and notes that English inputs perform at majority-vote level.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
HENNGE said on October 1 it set up HENNGE AI, a subsidiary with only two directors and no other employees, whe…

Writing for Robotics and Automation News, Unbox Robotics CEO Pramod Ghadge said reverse logistics needs AI dec…

GUGA said on October 1 it will significantly revise the syllabus for its 生成AIパスポート certification, applying it…

AMD announced AMD Ross, an agentic AI assistant for embedded design and development that runs across the AMD E…

DataSnipper CEO Vidya Peters said audit faces a staffing crisis, with two to two and a half times as many jobs…

Yann LeCun, the 2018 Turing Award winner, said he has "zero concerns" about rogue AI incidents like OpenAI age…
