
Opper published jevman, an open-source Pac-Man benchmark that runs six AI models — among them TypeSafe AI's Jev and GPT-6 Luna Decisions — for 100 games each, ranking them on average score, decision speed and cost. The longest game Opper recorded lasted 2 minutes 24 seconds.
Summaries like this, in your inbox every morning.
Jev, made by the US company TypeSafe AI, belongs to a different family from the large language models behind chatbots. Instead of writing prose, a probabilistic decision model takes supplied information and returns a probability for each of a preset list of options — for example, whether a customer inquiry is a return, a technical problem or something else. Because no explanatory text has to be generated, this style of model suits situations where a call has to be made fast.
The benchmark's rules reflect that emphasis. A game ends when all three lives are lost or 5 minutes pass, and the AI receives the maze layout, dot placement and ghost positions as JSON before choosing among up, down, left and right. When a model fails to answer inside its 2-second window, a simple fallback algorithm picks the direction instead, and those substitutions are counted — so a model's score reflects responsiveness as much as judgment.
The games are played against ghosts that follow the same rules as in the traditional Pac-Man, and a model that takes a long time to think its way to a correct answer may not finish near the top. Opper also set the benchmark up so several models, including Kev, Laya, Clef and GPT-6 Luna Decisions, can be compared inside the same game environment, and the official site lets visitors watch AI play or steer Pac-Man themselves against a model.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Adviser Itay Sagie points to Schneider Electric's roughly $22.6 billion equity deal for PTC, the Synopsys–Open…

Anthropic kicked off the Anthropic Cyber Mission, providing frontier Claude models, on-site engineers and thre…

Using OpenAI's Decisions API, model gpt-6-luna, a developer built a pipeline that answers questions like "is t…

A write-up describes running Claude Code unattended every four hours via Windows Task Scheduler and Python, wi…

The morning-news SKILL.md sat unchanged since June 17 at 2,116 bytes (version 1.2.0) even after its cron job g…

OpenAI said it disabled accounts behind Russia's "Dark Clark" operation, which used a fabricated researcher, M…
