AIToday
Large Language ModelsOpen-Source AIQiita 機械学習Published: Oct 9, 2026, 16:00 JST

TypeSafe AI's Jev clones: 14 open models tested on RTX 4090

TypeSafe AI's Jev clones: 14 open models tested on RTX 4090

On a self-built 76-question Jev-format set, six trained 3B–9B open-weight models (Imajev-4B, Clef-Flash 9B, JevK5-4B, Kev-4B, NeoHorse-Jev-4B, d1-3B) scored 88–92%, a spread under 3 questions, measured on an RTX 4090.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Jev's launch on 2026-09-15, alongside a $40M seed round, made TypeSafe AI's System One model a new category rather than just a product. The company's pitch is that chat models have become superhuman while automation has not kept up. Because Jev's weights, architecture and paper are all proprietary, the outside ecosystem could only copy its interface, not its internals.

That interface turned out to be easy to clone. Open-weight builders mostly took existing Open LLMs such as Qwen3.5 and Gemma 4 and swapped the generation head for something that reads out probabilities over choices. This article sorts those efforts into three methods: replacing the readout head, reading logits for lettered choices, and attaching a decision head to a bidirectional encoder. Each method carries a different trade-off between speed, flexibility, and reasoning depth.

Two independent benchmarks exist to sort the results. Benchmark Heaven's JevBench, which is unaffiliated with TypeSafe, scores models on a four-axis harmonic mean where cost is measured in USD per 1,000 judgments; Cloudflare runs a separate Decision Index, and the two do not agree on rankings. The author's own 76-question run is a reminder that most published model numbers are self-reported by their developers.

FAQ
What exactly does Jev output instead of text?
It returns a probability distribution over choices for each typed question. The three question types are noul (yes/no), choice (pick one from unordered options), and score (rate on an ordered rubric).
Do I need the Jev API to try this approach?
No. Open-weight alternatives like Kev let you reuse TypeSafe's Python SDK directly, and JevK5 has a GGUF version that runs on CPU via llama.cpp.
Can these open-weight models handle Japanese?
On the 9 Japanese questions in the test, 4B-class models and the multilingual Quyet-Small hit 88.9%, while the English-only 153M encoder scored 44.4%. TypeSafe itself notes Jev handles CJK languages but with lower accuracy.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGoogle unveils Gemini agent for cross-app work tasks