
What happened
The author served com-kotobalabs/open-jev-deberta-v3-large on Databricks Model Serving in AWS 東京リージョン, and on RTX PRO 6000 (GPU_LARGE_RTX) latency was 21.1/23.5 ms p50/p95 with throughput of 379 questions per second.
Why it matters
The endpoint was ready about 3 minutes after creation in Tokyo, so a company that cannot send data to an external API can keep decision-model inference inside its own Databricks boundary.
What to watch
Accuracy dropped to 0.691 on ood-test, questions the model was not trained on, so teams should evaluate on their own data before trusting new question types.
WHO IT HITSData and platform engineers who want to keep decision-model inference inside their own Databricks environment can now serve a Japanese-tuned model in Tokyo on RTX PRO 6000 rather than routing data to an external Jev API.
Summaries like this, in your inbox every morning.
Jev is a decision model: it takes a text and typed questions, such as multiple choice, a rating scale, or Yes/No, and returns answers with a probability distribution in a single inference pass. Because it does not generate text, the output format does not break, and it can handle tasks that people tend to ask an LLM to do, such as routing inquiries, judging attributes, and setting priorities, quickly and cheaply. The official Jev has not released its implementation and can only be used through an API.
The author picked com-kotobalabs/open-jev-deberta-v3-large from the openly available Jev-compatible models. The model card and the author's repository show that DeBERTa-v3-large, an encoder type, ran at 42 ms versus 676 ms for a LoRA-tuned LLaDA-MoE-7B, about 16 times faster, and a Qiita article reported accuracy of 95.92% on the Japanese sentiment classification task japanese_sentiment_classification (2,552 examples). The model is built on microsoft/deberta-v3-large with 435M parameters, under Apache-2.0, with a context length of 512 tokens and state limited to 256 tokens.
In the batch check, AI Runtime attached an A10G from the Tokyo workspace and reproduced numbers close to the model card: 0.861 accuracy on the 7,000-question test set versus 0.854 on the card, and 0.691 on the 8,000-question ood-test set versus 0.690 on the card. Because ECE, the calibration error, is small at 0.015 and 0.032, the confidence value can be treated as reliability as-is, for example by automatically handling high-confidence cases and having a person check low-confidence ones. How much of this workflow a given company can adopt is likely to depend on whether its own internal policy allows sending data to an external API in the first place.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
NetApp and Iterate.ai are packaging the AIPod Mini with Iterate's Generate platform and an embedded LLM, so en…
Mercor had 12 licensed CPAs work through simplified APEX Accounting Benchmark tasks

Testing Azure API Management's llm-token-limit policy at 800 tokens per hour, actual consumption hit 1,472 tok…

Qwen released Qwen3.8-Flash-Next on August 27, 2026, calling it a preview of the architecture planned for Qwen…

A Zenn article floated a hackathon where participants get the theme on the day, use no PC, internet, smartphon…

Anthropic's Message Batches API offers a 50% off rate, takes up to 10,000 requests per batch, and returns resu…
