AIToday
Large Language ModelsOpen-Source AIZenn AI/MLPublished: Oct 2, 2026, 22:00 JST

open-jev hits RTX PRO 6000: 21 ms, 379 q/s on Databricks

open-jev hits RTX PRO 6000: 21 ms, 379 q/s on Databricks

3 Key Points

  1. What happened

    The author served com-kotobalabs/open-jev-deberta-v3-large on Databricks Model Serving in AWS 東京リージョン, and on RTX PRO 6000 (GPU_LARGE_RTX) latency was 21.1/23.5 ms p50/p95 with throughput of 379 questions per second.

  2. Why it matters

    The endpoint was ready about 3 minutes after creation in Tokyo, so a company that cannot send data to an external API can keep decision-model inference inside its own Databricks boundary.

  3. What to watch

    Accuracy dropped to 0.691 on ood-test, questions the model was not trained on, so teams should evaluate on their own data before trusting new question types.

WHO IT HITSData and platform engineers who want to keep decision-model inference inside their own Databricks environment can now serve a Japanese-tuned model in Tokyo on RTX PRO 6000 rather than routing data to an external Jev API.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Jev is a decision model: it takes a text and typed questions, such as multiple choice, a rating scale, or Yes/No, and returns answers with a probability distribution in a single inference pass. Because it does not generate text, the output format does not break, and it can handle tasks that people tend to ask an LLM to do, such as routing inquiries, judging attributes, and setting priorities, quickly and cheaply. The official Jev has not released its implementation and can only be used through an API.

The author picked com-kotobalabs/open-jev-deberta-v3-large from the openly available Jev-compatible models. The model card and the author's repository show that DeBERTa-v3-large, an encoder type, ran at 42 ms versus 676 ms for a LoRA-tuned LLaDA-MoE-7B, about 16 times faster, and a Qiita article reported accuracy of 95.92% on the Japanese sentiment classification task japanese_sentiment_classification (2,552 examples). The model is built on microsoft/deberta-v3-large with 435M parameters, under Apache-2.0, with a context length of 512 tokens and state limited to 256 tokens.

In the batch check, AI Runtime attached an A10G from the Tokyo workspace and reproduced numbers close to the model card: 0.861 accuracy on the 7,000-question test set versus 0.854 on the card, and 0.691 on the 8,000-question ood-test set versus 0.690 on the card. Because ECE, the calibration error, is small at 0.015 and 0.032, the confidence value can be treated as reliability as-is, for example by automatically handling high-confidence cases and having a person check low-confidence ones. How much of this workflow a given company can adopt is likely to depend on whether its own internal policy allows sending data to an external API in the first place.

FAQ
What is a Jev decision model?
It takes a text plus typed questions such as choice, score, and Yes/No, and returns answers with probability distributions in one inference pass. Because it does not generate text, the output format does not break.
Which GPU should I use for open-jev on Databricks Model Serving?
RTX PRO 6000 delivered about 379 questions per second, versus 118 for A10G and 56 for T4. Accuracy was the same on all three GPUs, so the choice is about speed.
Can I use a GPU from the Tokyo region?
AI Runtime has no GPU resources in Tokyo yet, but enabling cross-region processing lets the Tokyo workspace attach an overseas GPU. In Model Serving, RTX PRO 6000 is now available in Tokyo.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleHao Global admits $160 million GPU smuggling to China