
What happened
On October 5, Liquid AI published the weights of its decision-focused model d1-3B on Hugging Face, and a Japanese benchmark against the d1 API found d1-3B scored 10.8 points lower on JNLI and 8.4 points lower on JCommonsenseQA.
Why it matters
The author cannot rule out that the API runs a larger model than d1-3B, so developers testing the open weights at home may not see the same accuracy they get from the API.
WHO IT HITSDevelopers and data scientists who plan to run d1-3B on their own GPU instead of paying for the d1 API may get lower accuracy on Japanese judgment tasks, so they may need to verify results against the API before relying on the open weights.
Summaries like this, in your inbox every morning.
Liquid AI published d1 as an API-only decision model, but the October 5 weight release of d1-3B on Hugging Face made it runnable on a single GPU such as an RTX 4080. To test whether the open version matches the paid one, the author ran both models on the same Japanese benchmarks: JNLI, JCommonsenseQA, and Livedoor news classification, feeding one item per request in the same order. The d1 API numbers came from an October 1 measurement, while d1-3B was measured on October 8, so the two runs are a week apart. d1-3B also required answer choices as a dictionary rather than a list, which was swapped in without explanation.
The paired comparison showed the API ahead on JNLI and JCommonsenseQA by margins that exceed chance, while the Livedoor result showed no meaningful difference. This pattern is consistent with the API running a stronger model than d1-3B, since the published weights only cover d1-3B and the smaller d1-omni-600M, and Liquid AI has not disclosed the size of the API model. The article's own limits qual_ify that comparison: it covers only three Japanese benchmarks, the API results were reused from an earlier date, and the open model received no explanation alongside its answer choices.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Common Sense Media's Youth AI Safety Institute tested ChatGPT for Teens with over 4,000 prompts and rated it "…

MSI Computer Japan began taking pre-orders on October 8 for the Prestige N16 Flip AI+ C1XM, a premium AI lapto…

OpenAI added GPT-6 and a new interactive interface, Intelligent UI, to ChatGPT's chat tab, starting a global r…

Microsoft announced MAI-Code-1.1-Flash, an AI model designed with local execution in mind, and proposed Hybrid…

OpenAI rolled out GPT-6 and Intelligent UI to Plus, Pro, Business, and Enterprise plans worldwide starting Oct…

SpaceX reportedly is arranging about $10 billion in bank loans and $30 billion in investment-grade bonds to bu…
