AIToday
Large Language ModelsAI Coding AssistantsVercel AI BlogPublished: Oct 6, 2026, 01:00 JST

Vercel's Python SDK adds evaluate() for Jev

Vercel's Python SDK adds evaluate() for Jev

3 Key Points

  1. What happened

    Vercel shipped the latest AI SDK for Python, installable with uv add ai, adding an experimental evaluate() API that talks directly to Jev via a model called typesafe-ai/jev.

  2. Why it matters

    Jev behaves like a universal classifier (it sorts inputs into categories without task-specific training), so developers can build narrow decision tools without collecting and labeling data per task.

  3. What to watch

    Vercel notes Jev will make mistakes, and its own tests showed gaps — it classified the partially typed Python expression "what's" + " up as English, for instance.

WHO IT HITSPython developers and data scientists who previously had to train a separate classifier for each narrow decision task can now test Jev through a single evaluate() call, though Vercel advises them to verify accuracy for their own use case.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Jev comes out of an older idea: the classifier, the classic primitive of machine learning. Traditionally you would train a classifier from scratch for each new task — collect labeled emails for spam filtering, for example, and tune the model's weights to that domain. The team behind Jev instead turned a large language model into a classifier, so it already carries broad human knowledge and does not need task-specific training. That is why the model is cheap and fast, and why it returns structured JSON that conforms to the types and choices you define rather than free-flowing text.

Vercel's Python team has wrapped Jev's API almost verbatim into the AI SDK, keeping the surface small — one evaluate() function plus a few supporting types. The blog walks through two hands-on tests. In the first, a Vercel engineer revisited a Python REPL feature that needed to tell English from Python; a hand-rolled classifier trained over two days felt jittery, and while Jev fared better, it still labeled the half-typed expression "what's" + " up as English. In the second, the engineer tried to make Jev write Python code by picking characters, then abstract syntax tree nodes; it produced syntactically valid but mostly incorrect code, and only after an LLM expanded the prompt into a detailed plan.

The takeaway is narrow but useful: Jev is built for tight decisions, not for general writing. Whether it is accurate enough for a given use case is something teams will have to test themselves — and the two examples suggest the answer depends heavily on how well the question can be pinned down to concrete choices.

FAQ
What is Jev and how is it different from a normal AI model?
Jev is a new kind of AI model that acts as a universal classifier. You feed it data and multiple-choice questions, and it responds with answers plus how confident it is in each — returning structured JSON rather than generated text.
What question types does the evaluate() API support?
It supports three: ChoiceQuestion (pick one answer), ScoreQuestion (rate on a scale you define), and NoulQuestion (estimate the probability a statement is true).
How do I install and start using it?
Install with uv add ai, then create an AI Gateway key and set AI_GATEWAY_API_KEY. The API takes a model, a state (a string or JSON), and a mapping of questions.
Vercel AI BlogRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGitHub launches ReviewBench, an open AI code review benchmark