
What happened
A developer ran TypeSafe's Jev, a non-generative judgment model on Cloudflare Workers AI, over 30 fictional care-worker job postings to flag 14 legally required labor conditions. Human labeling produced 97.8% agreement, 6.9% abstentions, and no critical false positives across 600 calls costing $0.05.
Why it matters
A model that returns only calibrated confidence can be trusted for narrow, high-stakes checks, as long as the most damaging error type is counted separately from overall accuracy, the dev's method shows.
What to watch
The thresholds that separate 'stated / ambiguous / not stated' must differ by state, and the build hinges on pinning the model version to jev-1.13.0 and capping calls, since Workers AI fallbacks silently change behavior.
WHO IT HITSCompliance and HR teams reviewing Japanese job postings, and any developer embedding AI checks into regulated workflows, gain a template for reliability: calibrated confidence, state-specific cutoffs, and human-verified labels.
Summaries like this, in your inbox every morning.
The test grew out of a specific compliance need: Japanese care-worker job postings must disclose 14 labor conditions under the Employment Security Act, and the developer built a checker that classifies each condition as stated, ambiguous, or missing. Rather than using one AI to grade another, the developer hand-labeled 30 sample postings — ten from Hello Work-style formats, ten from private job sites, and ten deliberately corrupted — and treated the most damaging error type separately. This mattered because a job listing that says 'not stated' when something is stated is the costliest mistake for this use case.
The other practical findings point to how such models should sit inside a real pipeline. Jev abstains when confidence falls below a threshold, and because the calibrated confidence behaves differently across states, the cutoffs were set per state and exposed in a config file so operations staff can adjust them. In production, the developer also pins the model version and caps calls per run, reflecting a concern that an unversioned third-party model could shift behavior without notice.
The broader pattern emerges from a second project: when classifying event announcements, Jev could reliably tell whether a text was a tournament notice, but not distinguish a new event from an alias or a moved one, because that requires matching against a historical ledger. That suggests judgement-only AI is likely best suited to narrow checks decided by a single document, with a rule-based system and human review handling anything that needs external records.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
ServiceNow launched Flow, a natural-language service desk that handles employee requests inside Slack and Micr…
DeepSeek and Huawei announced a partnership to develop semiconductor software, part of efforts to reduce China…

The AI Ataraxos beat Niemeijer, the most decorated Stratego player, with an 85 percent effective win rate over…

Google unveiled Gemini 4 Argon on September 30, saying it beats GPT-6 Astra and Claude Opus 5.5 on 13 of 19 be…

Yann LeCun told Fortune’s Emily Forlini he has zero concerns about rogue AI incidents, including OpenAI agents…

A Preply survey of over 5,000 professionals in nine countries found 92% of Gen Z respondents used AI for learn…
