AIToday
Large Language ModelsOpen-Source AIQiita 機械学習Published: Oct 8, 2026, 16:00 JST

Liquid AI's open d1-3B trails d1 API by 10.8 points on JNLI

Liquid AI's open d1-3B trails d1 API by 10.8 points on JNLI

3 Key Points

  1. What happened

    On October 5, Liquid AI published the weights of its decision-focused model d1-3B on Hugging Face, and a Japanese benchmark against the d1 API found d1-3B scored 10.8 points lower on JNLI and 8.4 points lower on JCommonsenseQA.

  2. Why it matters

    The author cannot rule out that the API runs a larger model than d1-3B, so developers testing the open weights at home may not see the same accuracy they get from the API.

WHO IT HITSDevelopers and data scientists who plan to run d1-3B on their own GPU instead of paying for the d1 API may get lower accuracy on Japanese judgment tasks, so they may need to verify results against the API before relying on the open weights.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Liquid AI published d1 as an API-only decision model, but the October 5 weight release of d1-3B on Hugging Face made it runnable on a single GPU such as an RTX 4080. To test whether the open version matches the paid one, the author ran both models on the same Japanese benchmarks: JNLI, JCommonsenseQA, and Livedoor news classification, feeding one item per request in the same order. The d1 API numbers came from an October 1 measurement, while d1-3B was measured on October 8, so the two runs are a week apart. d1-3B also required answer choices as a dictionary rather than a list, which was swapped in without explanation.

The paired comparison showed the API ahead on JNLI and JCommonsenseQA by margins that exceed chance, while the Livedoor result showed no meaningful difference. This pattern is consistent with the API running a stronger model than d1-3B, since the published weights only cover d1-3B and the smaller d1-omni-600M, and Liquid AI has not disclosed the size of the API model. The article's own limits qual_ify that comparison: it covers only three Japanese benchmarks, the API results were reused from an earlier date, and the open model received no explanation alongside its answer choices.

FAQ
How much lower is d1-3B's accuracy than the d1 API?
d1-3B scored 10.8 points lower on JNLI and 8.4 points lower on JCommonsenseQA, with only a 0.7-point difference on the Livedoor news classification task.
Are the differences statistically meaningful?
Yes. On JNLI, the API alone answered 68 items correctly versus 14 for d1-3B (p = 1.1e-9), and on JCommonsenseQA the split was 47 versus 5 (p = 1.3e-9). The Livedoor difference of 517 versus 461 was not significant (p = 0.08).
What could explain the gap?
The author lists three unverified possibilities: the API may run a larger model than d1-3B, it may run a different version of a same-size model, or the API may handle input such as answer choices differently.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGoogle opens SynthID Detector worldwide