AIToday
Large Language ModelsOpen-Source AIGIGAZINE AIPublished: Oct 9, 2026, 22:00 JST

Opper's jevman pits six AI models at Pac-Man, incl. GPT-6 Luna Decisions

Opper's jevman pits six AI models at Pac-Man, incl. GPT-6 Luna Decisions

Opper published jevman, an open-source Pac-Man benchmark that runs six AI models — among them TypeSafe AI's Jev and GPT-6 Luna Decisions — for 100 games each, ranking them on average score, decision speed and cost. The longest game Opper recorded lasted 2 minutes 24 seconds.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Jev, made by the US company TypeSafe AI, belongs to a different family from the large language models behind chatbots. Instead of writing prose, a probabilistic decision model takes supplied information and returns a probability for each of a preset list of options — for example, whether a customer inquiry is a return, a technical problem or something else. Because no explanatory text has to be generated, this style of model suits situations where a call has to be made fast.

The benchmark's rules reflect that emphasis. A game ends when all three lives are lost or 5 minutes pass, and the AI receives the maze layout, dot placement and ghost positions as JSON before choosing among up, down, left and right. When a model fails to answer inside its 2-second window, a simple fallback algorithm picks the direction instead, and those substitutions are counted — so a model's score reflects responsiveness as much as judgment.

The games are played against ghosts that follow the same rules as in the traditional Pac-Man, and a model that takes a long time to think its way to a correct answer may not finish near the top. Opper also set the benchmark up so several models, including Kev, Laya, Clef and GPT-6 Luna Decisions, can be compared inside the same game environment, and the official site lets visitors watch AI play or steer Pac-Man themselves against a model.

FAQ
What is jevman actually measuring?
It measures how quickly and reliably an AI picks a direction at each maze junction, based on JSON data about the maze, dots and ghost positions. Because each decision is capped at 2 seconds, raw reasoning power alone does not decide the score.
Can I run my own AI model on it?
Yes. jevman is open source, and any model that can answer an HTTP request can take part — cloud-hosted or running locally.
How is the ranking decided?
Models are ranked by their average score across 100 games, but a margin of error at a 95% confidence level is calculated as well, so models whose scores fall inside that margin are treated as tied. Each model's best single score is also shown.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleWinWay hits NT$10 billion revenue a year early on AI demand