AIToday
Large Language ModelsAI Coding AssistantsZenn AI/MLPublished: Oct 1, 2026, 22:00 JST

Tokyo dev validates typesafe/jev job-listing checks at 97.8% accuracy

Tokyo dev validates typesafe/jev job-listing checks at 97.8% accuracy

3 Key Points

  1. What happened

    A developer ran TypeSafe's Jev, a non-generative judgment model on Cloudflare Workers AI, over 30 fictional care-worker job postings to flag 14 legally required labor conditions. Human labeling produced 97.8% agreement, 6.9% abstentions, and no critical false positives across 600 calls costing $0.05.

  2. Why it matters

    A model that returns only calibrated confidence can be trusted for narrow, high-stakes checks, as long as the most damaging error type is counted separately from overall accuracy, the dev's method shows.

  3. What to watch

    The thresholds that separate 'stated / ambiguous / not stated' must differ by state, and the build hinges on pinning the model version to jev-1.13.0 and capping calls, since Workers AI fallbacks silently change behavior.

WHO IT HITSCompliance and HR teams reviewing Japanese job postings, and any developer embedding AI checks into regulated workflows, gain a template for reliability: calibrated confidence, state-specific cutoffs, and human-verified labels.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The test grew out of a specific compliance need: Japanese care-worker job postings must disclose 14 labor conditions under the Employment Security Act, and the developer built a checker that classifies each condition as stated, ambiguous, or missing. Rather than using one AI to grade another, the developer hand-labeled 30 sample postings — ten from Hello Work-style formats, ten from private job sites, and ten deliberately corrupted — and treated the most damaging error type separately. This mattered because a job listing that says 'not stated' when something is stated is the costliest mistake for this use case.

The other practical findings point to how such models should sit inside a real pipeline. Jev abstains when confidence falls below a threshold, and because the calibrated confidence behaves differently across states, the cutoffs were set per state and exposed in a config file so operations staff can adjust them. In production, the developer also pins the model version and caps calls per run, reflecting a concern that an unversioned third-party model could shift behavior without notice.

The broader pattern emerges from a second project: when classifying event announcements, Jev could reliably tell whether a text was a tournament notice, but not distinguish a new event from an alias or a moved one, because that requires matching against a historical ledger. That suggests judgement-only AI is likely best suited to narrow checks decided by a single document, with a rule-based system and human review handling anything that needs external records.

FAQ
How much did the Jev validation cost to run?
The developer measured $0.05 for 600 calls, processing roughly 1.22 million tokens. A single Jev call runs about $0.0001.
What questions is Jev good and bad at?
It works well when the answer is decided by reading one document, such as 'is this text a tournament notice?' It fails when the answer requires cross-referencing past data or a ledger, like distinguishing a new event from a renamed one.
How is personal data handled before sending job postings to Jev?
Names, phone numbers, and emails are replaced with [NAME], [TEL], and [MAIL] placeholders by regex before submission. Only replacement counts are logged, and the pasted text itself is not stored.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOllaya runs decision models locally, up to 255 options